Applications in Plant Sciences
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match Applications in Plant Sciences's content profile, based on 23 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Amoah, E. I.; Bunch, Z.; Thomas, H. M.; Patch, H. M.; Grozinger, C.
Show abstract
0.O_LIMorphological traits such as floral area and body size are fundamental to ecological research, serving as inputs for studies of pollinator-plant interactions, habitat quality, and biodiversity monitoring. However, accurately measuring these traits from images remains challenging, particularly in complex field conditions where existing tools exhibit reduced accuracy and limited generalizability across taxa. C_LIO_LIWe present EcoMorph, a modular morphological measurement system that leverages the Segment Anything Model 3 (SAM3) to quantify traits across diverse ecological contexts. Unlike task-specific segmentation models requiring domain-specific training data, SAM3s prompt-based architecture enables segmentation of arbitrary biological structures from natural-language prompts, using the same underlying model across flowers, insects, and other targets without retraining. From the resulting segmentations, EcoMorph extracts three classes of measurement: area, linear dimensions, and object counts. C_LIO_LIWe validated EcoMorph across two ecological scales. At the intermediate scale, EcoMorph-derived floral area agreed closely with manual ImageJ measurements (R2 = 0.935, n = 74) under simple-background conditions and (R2 = 0.928, n = 58) under complex-background conditions, with valid predictions for 95% of images. At the fine scale, EcoMorph-derived insect body area was strongly correlated with hand-measured intertegular distance (r = 0.810, n = 349), capturing body-size variation across species from the small Bombus impatiens to the large Xylocopa virginica. Object counts matched manual counts almost exactly for well-separated insects in an insect box (R2 = 0.9997, n = 12). C_LIO_LIBy combining prompt-based segmentation with modular measurement, EcoMorph enables high-throughput quantification of area, size, and abundance from heterogeneous image sources without taxon-specific training. This generality supports a broad range of ecological applications, including pollinator and plant trait research, biodiversity and abundance monitoring, and allometric biomass estimation. C_LI
Alves, R. T. d. L.; Gouvea, Y. F.; Dalapicolla, J.; Poczai, P.; Giacomin, L. L.
Show abstract
Premise: Genome skimming (GS) is a cost-effective approach for plant phylogenomics, but its ability to recover informative datasets from different genomic compartments, particularly genome-wide SNPs, remains poorly explored in Solanum. Methods: We evaluated shallow GS for phylogenetic inference in South American prickly Solanum lineages by recovering plastid, mitochondrial, and nuclear datasets, including coding regions and genome-wide SNPs. Phylogenies were inferred using maximum-likelihood and coalescent approaches under different SNP filtering strategies. Results: GS successfully recovered complete plastomes, organellar coding regions, and large SNP datasets, but failed to consistently assemble mitochondrial genomes or recover low-copy nuclear genes. SNP-based analyses, especially from the nuclear genome, produced stable, well-supported phylogenies that were largely congruent across inference methods. In contrast, coding-region datasets, particularly from the mitochondrial genome, showed greater topological discordance, revealing cytonuclear conflict. Discussion: Our results demonstrate that shallow GS is an effective strategy for generating informative SNP datasets for phylogenetic inference in Solanum, despite limitations in recovering complete mitochondrial genomes and low-copy nuclear loci. SNP-based analyses substantially expand the phylogenetic potential of GS, providing a practical and cost-effective alternative for systematic studies.
Kim, S.; Bowman, J.; Braun, E. L.; McDaniel, S.
Show abstract
Target enrichment sequencing using probe sets like GoFlag 408 has revolutionized phylogenetics, yet recent genomic data indicate that some probes may be sex-linked, potentially introducing topological conflict while also allowing studies of sex-specific evolutionary processes. To test for sex-linkage across the bryophytes, we developed UVfinder, a pipeline designed to identify sex-linked GoFlag loci across published moss genomes and enable sex-aware downstream analyses. Applying UVfinder to 50 dioicous moss genomes, we identified 93 probes that exhibit sex-linkage in one or more lineages, providing genomic evidence for neo-sex chromosome formation via autosome-sex chromosome fusion and gene translocation. Furthermore, by comparing species trees derived from sex-linked versus autosomal loci in Hypnales and Dicranidae, we demonstrate that sex-linked loci harbor phylogenetic information that is distinct from that in autosomes. We also discovered a pervasive female sampling bias in the genomic data, perhaps reflecting a preference among collectors for plants with sporophytes. Ultimately, our findings highlight the dynamism in sex linkage across bryophytes and suggest that sex-aware phylogenomics can be used to reconstruct ancestral karyotypes and potentially resolve topological conflict. We expect that UVfinder will facilitate the further study of sex-specific evolutionary processes, particularly with improved genome assemblies and increased sampling in males.
Gardette, A.; Belda, E.; Prifti, E.; Zucker, J.-D.
Show abstract
O_LIReference databases shape the taxonomic resolution, uncertainty, and reproducibility of metabarcoding analyses. For ITS barcodes, public references are distributed across repositories with different taxonomic conventions, geographic coverage, and annotation practices, creating conflicts, missing ranks, and misannotations when databases are merged or compared. C_LIO_LIWe introduce CurateMake, a reproducible Snakemake workflow for ITS reference database construction, harmonisation, and validation. It integrates four public sources (UNITE, BOLD, PLANiTS, and CALeDNA) and user-supplied databases, combines Catalogue of Life name harmonisation with ITSx-based region standardisation, MSA/HMM-based alignment grouping, and SATIVA phylogenetic validation. Raw, CoL-harmonised, and SATIVA-validated annotation layers are retained throughout to compare curation effects while preserving flagged records for review. C_LIO_LIWe evaluated CurateMake on 3.58 million ingested sequences and controlled error-injection simulations. ITSx expanded the final harmonised database to 5.19 million barcode-resolved entries by recovering ITS1 and ITS2 sub-regions from full-length ITS records. Across the full dataset, normalised intra-cluster entropy decreased from Raw to CoL-harmonised to SATIVA-validated annotations, consistent with improved taxonomic coherence. In simulations, CurateMake achieved the highest correction rate across 1%-50% corruption and, at 15% corruption, corrected 42% {+/-} 1% of introduced errors, compared with 28% {+/-} 1% for CoL alone and 0% for SATIVA without the workflows alignment infrastructure. C_LIO_LIThese results show that nomenclatural harmonisation and phylogeny-informed validation address complementary error classes, with phylogenetic validation contributing measurably only within taxon-coherent alignments in this benchmark. CurateMake therefore provides a reproducible, provenance-tracked framework for auditable ITS reference database curation in metabarcoding workflows. C_LI
Chen, C.-C.; Lehtonen, S.; Jefferson, P.; Fauskee, B.; Tuomisto, H.
Show abstract
Hybridization and introgression are thought to play key roles in the formation of biodiversity. However, detecting gene flow between species and understanding reticulation patterns remain challenging, especially in species-rich lineages with complex evolutionary histories. Adiantum is a large fern genus, and ecological studies in Amazonia have found that species identification is often difficult due to morphological similarity and overlapping characteristics among species. Although several hybrids have been described in tropical America, comprehensive studies investigating genetic exchanges in this genus are still lacking. We used chloroplast and nuclear phylogenomic data to examine evolutionary relationships among tropical American Adiantum species. By combining traditional phylogenetic analyses with advanced bioinformatic methods such as HybSeq-based target capture, reference-guided phasing (HybPhaser), and network-based analyses, we found widespread reticulate evolution involving both recent and ancient hybridization events. These were especially common among those species that have been difficult to delineate morphologically. Our findings indicate that reticulate evolution has played an important role in shaping the diversity of neotropical Adiantum, especially within the tetraphyllum lineage. Such widespread hybridization has no doubt contributed to morphological ambiguity and taxonomic challenges. Our integrated analytical approach provides a first attempt at untangling the reticulate evolutionary history in this group, and future studies with additional sampling can be expected to clarify the evolutionary processes further.
Zeng, Z.; Wang, Y.
Show abstract
Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.
Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.
Show abstract
Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.
Sims, B.;Gaudinier, A.;Blackman, B.
Show abstract
PremiseSeed size and morphology are critical traits in agriculture, ecology, and genetics, but high-throughput quantification of these traits is often limited by labor-intensive manual measurements or expensive, platform-specific imaging software. Methods and ResultsWe developed SeedMeasure, a lightweight, open-source, and cross-platform command-line tool written in Python that automates the measurement of seed area, length, and width from images. Using a simple imaging setup, the program processes images by correcting for perspective skew, filtering debris, and exports quantitative data alongside quality-check images. We validated SeedMeasure across nine diverse species, ranging from small Arabidopsis thaliana seeds to large Zea mays kernels. The tool quickly handles images using multithreading and demonstrates high reproducibility, yielding low coefficients of variation across repeated runs. ConclusionsCompared to existing software, SeedMeasure is free, offers faster processing through parallel computing, and provides standalone executables that require no programming dependencies. SeedMeasure offers an accessible, cost-effective, and high-throughput approach for rapid phenotypic profiling, making advanced seed morphological analysis available to researchers without specialized laboratory hardware.
Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.
Show abstract
Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).
Baldaszti, L.; Moonlight, P.; Brummitt, N.; Pironon, S.; Sarkinen, T.
Show abstract
Incomplete information on distributions for a high proportion of the world's plant species together with biases in global biodiversity data mean that current estimates of plant diversity patterns are skewed. A key issue is that current predictions rely on a subset of species that is not representative of all plant species. Here we tested the feasibility of a representative sampling approach for mapping global vascular plant diversity at the finest scale where comprehensive data is available. Using the World Checklist of Vascular Plants as a reference, we generate random samples of species with increasing sample sizes from the global species pool. We compare the diversity patterns retrieved from the samples against the patterns of the reference dataset using spatially weighted correlation coefficients and four different diversity metrics. We find that at the botanical country scale, representative global maps of species and phylogenetic diversity can be created with small numbers of species (~1% [0.2% and 0.4%, respectively]) at the botanical country scale. For effective growth form and family diversity sample sizes encompassing ~20% [19.2% and 19.5%, respectively] of all species are needed. Random samples require markedly fewer species to reach high correlations than when restricting the pool of species to single plant families or genera. We show that when representative samples are used robust inferences of plant diversity patterns can be made from only a small proportion of species.
Balaji, S.; Martinson, K. A.; Schellenberger, J. S.; Koley, J.; Inman, C. M.; Hofmann, H. A.; Young, R. L.; Harpak, A.
Show abstract
Biological research often requires information about species traits. Manual literature collation can be time-consuming and miss parts of the literature. To address this gap, we developed trAIt, a publicly available software for the retrieval of characteristics of species from scientific literature catalogued in the Europe PubMed Central (PubMed) database. trAIt provides a graphical user interface (GUI) in which users specify species and characteristics of interest. Leveraging a large language model (LLM), trAIt retrieves relevant papers, combines their content through a consensus-based summarization model, and outputs a species-by-characteristic table. For a case study involving frog species, trAIt recovered 47.1% of trait-species combinations in 2.75 hours, while an expert curator independently recovered 62.4% over months. The consensus-based summarization substantially aids accuracy compared to single-source extraction. Across three case studies of vertebrate taxa, an expert confirmed the accuracy of 70.9% of trait-species entries recovered by trAIt. We observed considerable variation across taxa in trAIts accuracy, which is possibly due to heterogeneity in open-access literature availability and inconsistencies in species and trait terminology. In sum, our analysis suggests that LLM-based tools can accelerate biological data synthesis but should be used to support domain experts research, rather than replace their judgment.
O'Brien, A.; Parada, P.
Show abstract
Deep-learning classifiers for the fungal internal transcribed spacer (ITS) report accuracies above 90% and are increasingly proposed for environmental metabarcoding. We benchmark two pretrained models, a convolutional network and a transformer sharing a training corpus of 5.23M sequences, against two established k-mer methods on 5,222 identical queries, one per genus, evaluated on both full-length ITS and the ITS2 subregion that dominates environmental sequencing. The design favours the classifiers: queries are drawn from the same public dataset they were trained on and stratified by whether a querys genus lies in their own label space, recovered from the distributed models, while the reference the k-mer methods consult mirrors that label space and excludes the queries themselves. Even so, on full-length ITS both classifiers are outperformed by both classical methods at every rank and in both strata: SINTAX recovers the correct family for 92.0% of seen-genus and 70.3% of novel-genus queries and best-hit alignment against a 56,327-sequence reference for 92.3% and 67.2%, against 77.9% and 57.7% for the transformer and 76.5% and 54.3% for the convolutional network. A hierarchical logistic regression on k-mer counts, fitted in ten minutes to 1.07% of the MycoAI training corpus, also exceeds both and places novel genera better than either search method, and refitted on ITS2 it recovers 89.2% of seen-genus families on that amplicon against 88.6% for best-hit alignment, so neither learned classification nor the amplicon is what fails. Restricting the identical records to ITS2 costs the k-mer methods 3.5 and 3.7 percentage points of seen-genus family accuracy but costs the classifiers 49.6 and 58.0, reducing them to 28.3% and 18.5%. An ablation identifies the cause. Grafting each querys unaltered ITS2 between the flanking regions of a donor record from a different phylum returns the donors family for 34.4% of queries against the querys own for 4.2% in the convolutional model, and 63.7% against 0.8% in the transformer, from a baseline of 0.1% where no donor sequence is present. The models therefore read taxonomy principally from the flanking regions rather than from the ITS2 barcode, which explains the collapse and predicts the same failure for any subregion amplicon. Compounding this, on ITS2 the classifiers output probability all but ceases to separate novel from known genera (AUROC 0.541 and 0.503, the latter at chance, against 0.866 for alignment identity and 0.785 for the SINTAX bootstrap), so the failure is not detectable from the models own output. We recommend that reported accuracies for such models specify the amplicon region of the evaluation, state the length distribution of the training corpus, and include a same-query classical baseline.
Mandelli, L.; Johnson, K. M.; Berretti, S.; Mencuccini, M.
Show abstract
1O_LIEmbolism, the formation of air bubbles in the plant water transport system, is a mechanistic driver of plant death. The Optical Vulnerability Technique (OVT) is an imaging method for non-invasive quantification of embolism (including P50, a common metric for drought vulnerability), which can also provide detailed spatial and temporal information. Its major cost lies in the post-processing of thousands of images. C_LIO_LIHere we designed, tested, trained, and make publicly available a neural network model to automate post-processing of OVT images. Using a dataset of 65 leaves from Senecio pterophorous, we compared our model predictions to results obtained via traditional post-processing by an expert. C_LIO_LIOur model resolved P50 to within 0.027 MPa of the expert-processed data with training taking 30 minutes to 2.5 hours and model-runtime in the order of seconds to minutes, demonstrating its promise for increasing the efficiency and throughput of P50 calculation. The models performance in replicating the pixels that constitute embolism events was lower (mean event-frame IoU of 0.38). C_LIO_LIWe invite the community to utilise our model but emphasise that it does not replace the expert-processing pipeline and that care must be taken when considering applying this and similar approaches to OVT data. C_LI
Kirschke, G. E.; Bain, J. A.; Ogilvie, J. E.; CaraDonna, P. J.
Show abstract
O_LIFloral nectar plays a critical role in shaping the ecology and evolution of plant-pollinator interactions. Effective and efficient methods that allow for broad-scale sampling of nectar volume and sugar concentration across a diversity of taxa are needed to improve our understanding of many dimensions of mutualistic plant-pollinator interactions--including their basic ecology and evolution, their responses to environmental change, and their conservation and restoration. C_LIO_LIDespite the key importance of nectar for mediating plant-pollinator interactions, quantifying floral nectar in the field from many different plant species is challenging because there is often no one-size-fits-all sampling method that is effective across a diversity of floral structures and nectar traits. Different methods require different preparation, and sampling from many species involves a variety of logistical challenges. C_LIO_LIHere we provide a methodological roadmap for sampling floral nectar in the field from many different plant species. We describe our nectar collection methods in detail, including necessary equipment, calculations, and approaches appropriate for different floral morphologies. We also provide a troubleshooting guide for common problems encountered while collecting nectar in the field. To demonstrate the utility and effectiveness of our methods for collecting nectar from many different species, we present results on nectar trait variation from 53 species in an ecosystem. C_LIO_LIOur method illustrates that nectar traits vary considerably within and among plant species, indicating that large-scale nectar sampling projects are an important consideration for many basic and applied questions in pollination ecology and evolution. We hope that across many plant communities and ecosystems, our paper provides a practical roadmap for how to navigate the complexities of quantifying floral nectar traits. C_LI
Trauden, T.; Rakotomalala, A. A. N. A.; Junker, R. R.; Sauressig, L.; Trauden, K.; Munoz Andres, M.; Dannoritzer, R.; Farwig, N.; Pinkert, S.
Show abstract
Leaf shape is a fundamental trait of plant ecological strategies, influencing biotic interactions and ecosystem functioning. However, established quantitative metrics fail to capture subtle variations and irregularities, require user-based reference points or are challenging to compare among taxa with broadly different leaf shapes. In addition, established metrics typically conflate (aggregate) leaf edge complexity and macro-shape complexity, despite their independent functional significance and genetic foundations. Here, we introduce an entropy-based framework to quantify two new complexity metrics: edge complexity and macro-shape complexity. Based on three case studies, we show that these metrics outperform aggregate metrics in predicting Quercus robur chemical traits, provide more intuitive interspecific classifications, and strongly align with human perception. In addition, edge and macro-shape complexity show high complementarity, while aggregate metrics are highly redundant and typically strongly related to leaf area. Emerging as the strongest predictor of leaf chemistry and key visual cue for complexity as perceived by humans, the effects of edge complexity highlight the under-appreciated functional significance of leaf margins. Our framework and the proposed entropy-based complexity metrics thus promise to help unlock the potential of growing digital image archives of leaves, including images from herbaria and fossils, and are technically readily applicable to shapes of algae, bacteria, pollen, and beyond. The accompanying package ShapeComplexity enables the broad application of entropy-based metrics, providing a powerful tool to explore how the shape of organisms and biological structures influences ecological strategies, biotic interactions, and ecosystem functioning while tracking spatial and temporal variation.
Slimp, M.; Martinez, L. N.; Kapp, J. D.; Kirby, M. E.; MacDonald, G.; Hankins, D. L.; Melrose, S.; Johnson, M. G.; Shapiro, B.; Meyer, R. S.
Show abstract
As we face the sixth mass extinction, understanding how ecosystems have persisted--or collapsed--through millennia of changing climates and human activity is critical for preventing biodiversity loss. We bolstered the past 24,000 years of plant and mammal records using targeted capture of ancient sedimentary DNA (sedaDNA) from Southern Californias Lake Elsinore, a cultural center for the Payomkawichum (Luiseno), Cahuilla, and other Peoples. Our sedaDNA approach generated a diverse dataset that included 18 plant orders not previously documented from Lake Elsinore. We paired these records with local measurements and paleo evidence of fire regimes, climate, demographic history, and ethnobotanical knowledge. We find that ecological stability persisted for 10,000 years of continuous human presence, reflecting ecosystem resilience through major climatic shifts, altered fire regimes, and varying intensities of Indigenous land use. SedaDNA revealed increased availability of food, medicinal, and utilitarian plant taxa during this period of botanical stability, shedding light on ancient fire-environment-human interactions that can inform contemporary management strategies.
Maciel, E. A.
Show abstract
Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.
Wu, T.; Yang, Z.; Shi, J.; Zou, M.; Wu, Y.; Jiang, S.; Xia, C.; Kong, L.; Yang, L.; Xia, Z.
Show abstract
Plant functional genomics requires the integration of sequence, expression, evolutionary, regulatory and literature evidence. However, the corresponding analyses are often distributed across disparate programs, scripts and databases, creating substantial barriers to task organization and result interpretation. Here, we present PlantAI, a multi-agent system that integrates bioinformatics analysis, project-level process tracking and knowledge-assisted interpretation. A Main Agent coordinates two complementary routes: an analysis route that invokes bioinformatics tools for RNA-seq and gene-family analyses, and a knowledge route that uses PlantAI-RAG for knowledge retrieval and evidence synthesis. PlantAI-RAG currently contains 31,207 plant-science literature records, comprising approximately 3.82 million normalized entities and 8.25 million literature-supported relation assertions. In an evaluation using plant-science questions, it achieved a Gold evidence-assertion recall of 86.7%, while strict accuracy ranged from 77% to 82% across three independent evaluator models. We further demonstrate an end-to-end task using 24 rice RNA-seq libraries collected under salt stress, spanning transcriptome analysis, candidate-family screening, HXK/HKL family analysis and knowledge-assisted interpretation, and prioritize OsHXK8 for experimental validation. By preserving analysis artifacts, run manifests, logs and environment records, PlantAI supports result verification and repeat execution while linking project-derived results to traceable literature evidence. Together, these capabilities provide an integrated and auditable framework to support plant functional genomics research.
Gatula, L.; Bezrukov, I.; Atemia, J.; Chapano, C.; Gamundani, P. T.; Zimudzi, C.; Chatukuta, P.
Show abstract
Herbaria serve as invaluable spatio-temporal repositories of plant diversity information. Digitization of herbarium collections enhances the accessibility, discoverability, and long-term preservation of this important plant information, yet financial and infrastructural constraints often prevent herbaria in resource-constrained regions from digitizing their collections. Consequently, critical plant diversity data gaps remain due to underrepresentation of these collections in global biodiversity databases. Here, we describe an AI-assisted modular digitization toolkit specifically designed for herbaria operating under limited funding, developed and refined through our experience digitizing the crop wild relative (CWR) collection of the National Herbarium of Zimbabwe. The toolkit comprises three core components: (1) a portable, cost-effective photostation assembled from commodity parts, (2) a streamlined cascade workflow for systematic digital imaging, and (3) an AI-assisted data management pipeline for image quality control, label transcription, data analysis, and presentation. Compared to manual transcription and legacy optical character recognition approaches, AI-based transcription achieves lower time cost while maintaining high accuracy, and AI-driven data management delivers accessibility and reduced expenditure relative to conventional database infrastructure. The toolkit is designed to allow herbarium staff full autonomy over the digitization procedure, ensuring institutional ownership and the capacity for independent continuation beyond initial project support. By prioritizing affordability, modularity, and simplicity, this toolkit provides a replicable framework that may enable resource-constrained herbaria to locally generate high-quality scientific data for conservation and the sustainable utilization of plant genetic resources.
Baldwin, B. G.; Fawcett, S.; Kostic, A.; Wessa, B. L.; Freyman, W. A.
Show abstract
Premise of the studyDisentangling ancestry of insular plant lineages is important for resolving their evolutionary histories but often complicated by polyploidy, an overrepresented condition in remote island floras. A phylogenomic approach was used to study the origin of subgenomes of the tetraploid Hawaiian silversword alliance and advance understanding of this major example of adaptive radiation. MethodsSequences from 461 target-capture nuclear loci were phased to subgenomes of the tetraploid silversword alliance (Compositae) using assemblies from diploid North American tarweeds as clade references. Maximum-likelihood and multispecies- coalescent analyses of phased loci were conducted with and without representatives of diploid lineages throughout tribe Madieae. Key resultsAn allopolyploid origin of the silversword alliance was corroborated, with Subgenome-A phased consistently to Anisocarpus madioides and Subgenome-B phased to diploid representatives of each of three other sublineages of the "Madia" lineage. Phylogenomic analyses indicate Anisocarpus madioides is more closely related to Subgenome-A than to other tarweeds, including Anisocarpus scabridus. Subgenome trees for the silversword alliance are highly congruent with each other and with structural genomic differences between lineages. ConclusionsThese results indicate an even closer relationship of the silversword alliance to extant California tarweeds than previously understood. Anisocarpus madioides has chromosomal and ecological traits anticipated in an ancestor of the silversword alliance, with the widest geographic range among perennial tarweeds and sticky bracts completely enveloping fruits. Within the silversword alliance, robust support for four major clades reinforces the phylogenetic value of chromosomal interchanges and taxonomic insights by Asa Gray regarding the enigmatic liana Dubautia latifolia.